OpenAI just handed us the scorecard. We were built to pass it.
Schema Driven · July 2026 · 8 min read
On July 17th, OpenAI's CFO published a scorecard for the AI age. Read past the product placements and it's the most useful thing a frontier lab has said to a buyer in a year: it tells every CFO in the market to stop measuring AI by cost per token and start measuring the full cost of a successful outcome.
That's not their framework. That's ours. And two weeks earlier, Palantir's CEO had torched the same meter from the other side on live TV. The two loudest voices in enterprise AI just bracketed our entire pitch, then wrote it on official letterhead and mailed it to every prospect we have.
"A lower-cost model may have cheaper tokens, but getting great results may require more attempts, more time, or more human review."
Sarah Friar, CFO, OpenAI
That is the Tokenomics argument, verbatim, from the vendor whose tokens you'd be renting. When the category leader concedes that the cheap meter is the expensive one, the conversation is no longer whether outcomes are the unit. It's who actually delivers them.
The scorecard has four questions. We have four answers.
The scorecard boils down to one metric, call it useful work per dollar, behind four questions. Here is each question as OpenAI poses it, and what our architecture does about it.
Is AI completing work that matters, and how much?
Their "done" is a workflow finished in the system where the work lives. That's the whole point of capturing a domain as a validated schema: we generate the boring, correctness-critical 80% deterministically, so the model spends its budget on the judgment 20% that actually moves the outcome. Work shipped, not drafts produced.
What does one successful outcome actually cost?
Full cost ÷ tasks that met the bar, retries and review and rework included. A deterministic core has a cost per task of essentially zero and a success rate of one: it doesn't re-reason an answer with a fixed shape. The meter only runs on the part that genuinely needs a model. Fewer attempts by construction.
How often does the work come back right?
She splits results into ready to use, needs correction, and needs escalation. Business rules as tolerances, a judge at every station, and a design system as the answer key push work into the ready bucket and stop the wrong shape from ever leaving the line. Dependability is the deliverable.
Does each dollar buy more work over time?
Her answer is compute economics. Ours is ownership: every rule you validate widens the deterministic core, and signed deltas from the patent-pending Mempack™ process sharpen the experts without a retrain. Next year's regeneration is free where this year's cost money. Spend converts into an asset you keep.
Two weeks earlier, the other side said the quiet part
OpenAI's CFO didn't invent this framing in a vacuum. On July 1st, Palantir's Alex Karp went on CNBC's Squawk Box and torched the same meter from the opposite direction. Where the CFO's memo is polite about it, Karp is not.
"The basic view among enterprises in this country is I'm going to chillax and waste my time with tokens, I'm gonna get no value, and they're gonna get my IP."
Alex Karp, CEO, Palantir · CNBC, July 1, 2026
He described customers moving away from "tokenmaxxing" toward ROI, and the word he kept returning to was ownership: enterprises wanting to know they "own the means of production," that it isn't being transferred to someone else. His sharpest fear is the one no scorecard measures. That a lab will "take the alpha of my business, transfer in their weights, and compete against me."
So here's the pincer. The category leader's CFO says stop paying for tokens, pay for outcomes. The loudest applications CEO says the token model quietly walks off with your IP. One indicts the meter; the other indicts the ownership transfer behind it. They're describing two faces of the same problem, and between them they've drawn the exact shape of what we sell.
The one line neither of them can offer, and we can
The scorecard is honest about its four measures. It's quiet about a fifth, the one Karp keeps hammering: what do you own when the engagement ends? Run the scorecard on a pure token relationship and even a perfect score leaves you with prompts, brittle integrations, and a renewal. The intelligence, the weights, and your accumulated alpha stay on someone else's meter. Their framework optimizes the rental. It has no row for the deed.
That's the row we win outright, and it's the one Karp says enterprises actually care about. Same four questions, but the compounding asset (the specs, the generated apps, the domain experts) is yours to run, fork, and keep. No one transfers in their weights and competes with you on your own data, because the weights and the data are yours. The scorecard tells you to measure useful work per dollar. Ownership is what happens to every dollar after the measurement.
So we're adopting it
There's no point fighting a framework the whole market is about to use, especially one that indicts the token business and describes ours. So we're taking the scorecard as our own rubric. If a prospect wants to grade us on useful work, cost per successful task, dependability, and return at scale, we'll hand them the pencil. We built the thing to pass this exact test; someone else just made it the standard.
Bring your own workflow and we'll score it together: full cost, success rate, and what you hold at the end. Request access →
Quotations are from OpenAI's July 17, 2026 post (linked above) and Alex Karp's July 1, 2026 CNBC Squawk Box appearance, as reported. The four-question framing is OpenAI's; the ownership critique is Palantir's; the architecture and the fifth question are ours.